Skip to main content

Selecting and Transforming Series

Select without ambiguity​

import pandas as pd

s = pd.Series([90, 80, 70], index=["Ada", "Lin", "Sam"], dtype="float64")
s.loc["Ada"] # one label
s.loc[["Ada", "Sam"]] # several labels
s.iloc[0] # first position
s.iloc[:3] # first three positions
s[s.ge(80)] # boolean filter

Use .loc for labels and .iloc for positions, especially when the index itself contains integers. Bare s[key] is concise for an unambiguous label, but explicit indexers communicate intent and survive index changes better.

Transform in this order​

  1. Use native vectorized arithmetic or string/datetime accessors.
  2. Use map for a scalar lookup or elementwise function.
  3. Use where, mask, replace, or fillna for conditional replacement.
  4. Use apply only when no clearer vectorized operation exists.
normalized = (s - s.mean()) / s.std()
bands = s.map({90: "A", 80: "B"})
clipped = s.clip(lower=0)

A dictionary passed to map turns unmatched values into missing values. Use replace when unmatched values should remain unchanged.

Combine Series​

Use pd.concat to stack or place objects side by side:

train, test = s.iloc[:2], s.iloc[2:]
actual, predicted = s, s + 5
stacked = pd.concat([train, test], ignore_index=True)
table = pd.concat([actual.rename("actual"), predicted.rename("predicted")], axis=1)

Series.append is not the combination API. Before column-wise concatenation, check whether label alignment is intended.

Iteration​

If iteration is genuinely necessary, s.items() yields (label, value) pairs. Do not use iteration for arithmetic or routine filtering; vectorized expressions are clearer and usually faster.

Results and statistical assumptions​

For this input, s.loc["Ada"] is the scalar 90.0, selecting two labels returns a length-2 Series, and the filter retains Ada and Lin. The transformations return new Series; s is unchanged. bands contains "A", "B", then a missing value: this is an exact lookup, not score-range binning.

Series.std uses the sample denominator (ddof=1), so the mean is 80, standard deviation 10, and normalized is [1.0, 0.0, -1.0]. Use ddof=0 for a population standard deviation. A constant Series has zero deviation; fewer than two valid values cannot define the sample deviation. Decide how to handle these cases before dividing.

Source​

Explore connectionsOpen network